Papers with building process
Going “Deeper”: Structured Sememe Prediction via Transformer with Tree Attention (2022.findings-acl)
Copied to clipboard
| Challenge: | Existing studies ignore hierarchical structures of sememes in sememe-based semantic description systems. |
| Approach: | They propose a structured sememe prediction problem to predict a sememes tree with hierarchical structures rather than a set of sememas. |
| Outcome: | The proposed model outperforms baseline models and shows its effectiveness . it predicts a sememe tree with hierarchical structures rather than a set of sememes . |
WordNet-Shp: Towards the Building of a Lexical Database for a Peruvian Minority Language (L18-1)
Copied to clipboard
| Challenge: | WordNet-like resources are lexical databases with highly relevance information and data that could be exploited in more complex computational linguistics research and applications. |
| Approach: | They propose to build a WordNet database for a low-resourced and indigenous language in Peru . they propose to use word2vec similarity to compare definition glosses in a dictionary with the content of a Spanish WordNet . |
| Outcome: | The proposed database is based on a bilingual dictionary written in Spanish and an automatic evaluation process using a manually annotated Gold Standard in Shipibo-Koniba. |
TArC: Incrementally and Semi-Automatically Collecting a Tunisian Arabish Corpus (2020.lrec-1)
Copied to clipboard
| Challenge: | Arabish is a spontaneous coding of Arabic dialects in Latin characters and "arithmographs" this code-system was developed by Arabic-speaking users of social media . little research has been dedicated to Tunisian Arabish (TA) |
| Approach: | They describe the constitution process of the first morpho-syntactically annotated Tunisian Arabish Corpus . they describe preliminary work on the TArC semi-automatic construction process . |
| Outcome: | The first morpho-syntactically annotated Tunisian Arabish corpus (TArC) was developed by arab-speaking users of social media . the code-system will be a useful support for different types of analyses, computational and linguistic, as well as for NLP tools training. |
French Tweet Corpus for Automatic Stance Detection (2020.lrec-1)
Copied to clipboard
| Challenge: | a new corpus of tweets is being developed for automatic stance detection of fake news . the task involves determining the attitude expressed in a text toward a target . this is a difficult task to overcome as discussions about fake news are controversial . |
| Approach: | They propose to build a human-annotated corpus for automatic stance detection of tweets in french . they propose to use four classes broadly adopted by the community for annotation . |
| Outcome: | The proposed corpus is the first freely available stance annotated tweet corpus in the french language. |